Skip to content

Deduplicate test fixtures and prune subsumed tests - #1163

Merged
robinebers merged 1 commit into
mainfrom
claude/audit-test-suites-759d8e
Aug 25, 2026
Merged

Deduplicate test fixtures and prune subsumed tests#1163
robinebers merged 1 commit into
mainfrom
claude/audit-test-suites-759d8e

Conversation

@robinebers

@robinebers robinebers commented Aug 25, 2026

Copy link
Copy Markdown
Owner

TL;DR — Second test-simplification pass after #1143, focused on the suites that pass didn't touch: consolidates duplicated test fixtures, parameterizes near-identical test clusters, and removes tests whose behavior is already verified elsewhere. Net −360 lines (724 deleted, 364 added) with the full suite still green.

What was happening

  • Four test files each carried their own private copy of the Cursor JWT builder and an identical key/value SQLite fake; the OpenCode scanner and provider suites duplicated their SQLite stub and row builder; BlockingParser existed twice (integer and Claude-entry variants); and TogglingProviderRuntime / a private SequenceProviderRuntime were the same idea implemented twice.
  • Several suites repeated ~25-line provider constructions per test (ClaudeDesktopAuthStore, Antigravity error-category tests).
  • A handful of tests were near-identical clones (four Grok Bot optional-endpoint tests, Grok credits absent/non-numeric field pairs, Devin credentials-TOML pair), byte-identical duplicates (a ReorderGeometry test that only renamed a string ID), or restated the implementation in the test loop (two ShareCard condensed-row tests recomputed the production rule instead of asserting expected output).
  • A few single tests were fully subsumed by stronger tests elsewhere (PanelHeightCoordinator, AntigravityLayout pins, ZAI empty-limits, GrokCreditsConfig monthly period, MenuBarContent compaction, LocalUsageAPI limits-404, OpenCode zen-only scan, ShellEnvironmentSnapshot roundtrip, ClaudeDesktopAuthStore revoked-CLI fallback pair).

What this changes

  • Adds shared makeCursorJWT and KeyValueSQLite to TestSupport.swift; the four Cursor suites now use them.
  • Shares the OpenCode SQLite stub (OpenCodeFakeSQLite) and openCodeRow builder between the scanner and provider suites.
  • Replaces TogglingProviderRuntime with one shared SequenceProviderRuntime (different snapshot per refresh, call-counting); makes BlockingParser generic over the item type.
  • Extracts makeProvider/makeAuthStore helpers in ClaudeDesktopAuthStoreTests and a makeCloudCodeProvider helper in AntigravityProviderTests, collapsing repeated construction boilerplate.
  • Parameterizes the four Grok Bot optional-endpoint tests into one table-driven test (all four cases kept), and merges the Grok credits, Devin TOML, and OpenCode hasLocalCredentials pairs.
  • Removes duplicated or subsumed tests listed above, each verified as covered elsewhere before deletion; the ShareCard pair is replaced by one fixed-expectation test on a hand-built fixture, and MenuBarContent's unique $130 rounding case moved into MetricFormatterTests where the rule lives.

Heads-up

  • Deliberately untouched: LayoutStoreTests and the large Claude/Codex/Copilot suites (every test pins a distinct behavior or documented regression), LogRedactionTests' verbatim Rust-parity cases, all boundary/clamp tests, and every concurrency/ownership/account-isolation test.
  • No production code changed — all 25 modified files are under Tests/.

Tests

  • swift build --build-tests — clean (only pre-existing warnings).
  • swift test — 1,173 tests passed, 0 failures, 3 skipped (pre-existing env-gated live/parity tests), plus the CLI and swift-testing bundles. Baseline run before the changes was also green.

🤖 Generated with Claude Code


Note

Low Risk
Only test code under Tests/OpenUsageTests is modified; behavior is unchanged and risk is limited to possible loss of unique test coverage, which the PR explicitly addresses by consolidating rather than deleting scenarios.

Overview
This PR is a test-only cleanup (~−360 lines): no production code changes.

Shared fixturesmakeCursorJWT, KeyValueSQLite, SequenceProviderRuntime (replacing TogglingProviderRuntime), and a generic BlockingParser land in TestSupport.swift / JSONLScannerTestSupport.swift. Cursor suites drop four duplicate JWT/SQLite copies; OpenCode shares openCodeRow and OpenCodeFakeSQLite between scanner and provider tests.

Less boilerplatemakeCloudCodeProvider in Antigravity provider tests and makeAuthStore / makeProvider / cliCredentials in Claude desktop auth tests replace repeated ~25-line setups.

Consolidated cases — Four Grok Bot optional-endpoint tests become one table-driven test; Grok credits absent-field tests, Devin HTTPS-vs-HTTP TOML tests, and OpenCode hasLocalCredentials positive/empty paths are merged. ShareCard condensing is asserted on a fixed hand-built fixture instead of re-implementing the production rule in the test.

Removed redundancy — Tests dropped where coverage already exists (e.g. Antigravity pins exact-set, MenuBar compaction → MetricFormatterTests for $130 tray rounding, duplicate Claude revoked-CLI fallback, LocalUsageAPI limits 404, and others noted in the PR description).

The full suite remains green (1,173 tests).

Reviewed by Cursor Bugbot for commit d0771a3. Bugbot is set up for automated code reviews on this repo. Configure here.

Second simplification pass after #1143, covering the suites that pass
did not touch: consolidates duplicated fixtures (Cursor JWT/SQLite
fakes, OpenCode stub, BlockingParser, sequence runtimes), parameterizes
near-identical test clusters, and removes tests whose behavior is
verified elsewhere or that restated the implementation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions github-actions Bot added the tests label Aug 25, 2026
@robinebers
robinebers merged commit a7f603e into main Aug 25, 2026
3 checks passed
@robinebers
robinebers deleted the claude/audit-test-suites-759d8e branch August 25, 2026 04:25
liaomingxin added a commit to liaomingxin/openusage that referenced this pull request Aug 26, 2026
Brings in the 8 upstream commits behind v0.7.10-beta.3: multi-account Claude
(robinebers#1164), sub-1% session countdown (robinebers#1167, WidgetData.isSessionWindow →
sessionStartSignal), Codex Session on by default (robinebers#1165), GLM 5.3 pricing
(robinebers#1171), test-fixture dedup (robinebers#1163), plus docs/policy chores.

Conflicts and how they were resolved (both sides kept everywhere):

- Sources/OpenUsage/Services/ProviderAccountAssembly.swift — kept both the
  fork's Codex extra cards (CodexExtraCard, extraCodexCards, the
  ~/.cli-proxy-api discovery, mergedObservations/mergedSources) and upstream's
  Claude Desktop organization discovery (ClaudeAccountCard, claudeCards,
  allowsUnattributedPiUsage, the bare "claude" identity key rewritten to the
  org record id). Restructured the core make() so it appends the extra-Codex
  observations, then runs the Claude desktop-org discovery, then reconciles
  exactly once per pass through mergedObservations, then builds BOTH card
  lists — upstream's two early returns are now conditions on the discovery
  block, so a skipped/legacy Claude path still emits the Codex extra cards.
  Kept the fork's cold-login-shell guard (!families.isEmpty ||
  !extraCodex.isEmpty) and the extra-Codex read in
  make(defaults:waitsForLoginShell:). init takes both extraCodexCards and
  claudeCards with defaults so upstream and fork tests both compile. 361 LOC,
  no split needed.

- Sources/OpenUsage/Providers/ProviderCatalog.swift — one
  make(defaults:extraCodexCards:claudeCards:claudeIdentityKeys:): upstream's
  Claude card list first, then CodexProvider() followed by the fork's extra
  Codex cards, then Cursor and the alphabetical tail with Kimi after Grok.
  Comment updated to mention both.

- Sources/OpenUsage/App/AppContainer.swift — call site passes all three
  arguments (extraCodexCards, claudeCards, claudeIdentityKeys).

- Sources/OpenUsage/Services/UsageReader.swift — same three arguments; kept
  upstream's placement of ProviderEnablementStore(defaults:) after the
  registry is built.

- docs/menu-bar.md — default-star sentence lists both upstream's Codex
  Session/Weekly and the fork's Kimi Session/Weekly.

Sanity-checked DefaultLayout.swift (auto-merged): kimi rows still in
metricIDs/pinnedMetricIDs/expandedMetricIDs, codex.session now in metricIDs
and pinnedMetricIDs.

Verified: swift build and swift build --build-tests both succeed;
ProviderAccountAssembly / CodexMultiAccount / ClaudeDesktopAuthStore /
CodexExtraCardCatalog tests pass (24/24).
mstallone added a commit to mstallone/runway that referenced this pull request Sep 5, 2026
## TL;DR

Selective re-implementation of the OpenUsage commits since Runway #111
that still apply: Claude Desktop's account-prefixed token caches, Claude
spend tiles without an OAuth login, Grok subagent session ledgers, Codex
Business Premium, Cursor's Models/Other Models labels, and new model
rates.

## What was happening

- Upstream has moved on since #111. Each new OpenUsage commit was
reviewed against Runway's architecture, existing ports, and the "don't
bloat" bar.
- Recent Claude Desktop builds store tokens under `acct:<user>|<legacy
key>` (openusage robinebers#1212). Runway expected the client UUID first and
skipped those entries, so a Desktop-only login showed Not logged in.
- Claude returned a hard authentication error before scanning local logs
when no OAuth login existed (openusage robinebers#1138), so API-key gateway users
lost Today/Yesterday/Last 30 Days.
- Grok's scanner skipped every `subagent*` session (openusage robinebers#1193).
Child work that the coordinator turn did not include disappeared from
spend.
- Codex's `self_serve_business_prolite` entitlement rendered as "Self
Serve Business Prolite" (openusage robinebers#1194).
- Cursor's dashboard now calls the two model pools **Cursor Models** and
**Other Models**; Runway still said Auto Usage / API Usage (openusage
robinebers#1134, labels only).
- GPT-6 Astra, Gemini 3.8 Flash, Fable 5.1, GLM 5.3, and Grok Bot CSV
slugs had no supplement entries, so those rows tripped the
unpriced-model warning.

## What this changes

- Desktop cache selection strips the `acct:<user>|` prefix, keeps only
the signed-in account (from `lastKnownAccountUuid`), and lets a scoped
tombstone suppress the matching legacy V1 alias.
- Unauthenticated Claude refreshes still scan local logs. Spend tiles
render under the existing Not logged in notice when those logs contain
usage; an empty machine stays a hard error card.
- Grok scans every durable `updates.jsonl` ledger. Prompt-id dedup still
drops forked parent replays. `summary.json` is no longer required to
keep a ledger.
- Codex maps `self_serve_business_prolite` to **Business Premium**.
- Cursor widget IDs are unchanged. Titles/labels become Cursor Models /
Other Models to match Cursor's dashboard.
- Pricing supplement: GPT-6 Astra (OpenAI card, 2× fast), Gemini 3.8
Flash (Cursor table, $3.50 output), Fable 5.1, GLM 5.3, and `grok-bot-*`
→ Grok 4.6.

## Heads-up

Reviewed and **not** ported, with reasons:

- **openusage robinebers#1116 / robinebers#1127 / robinebers#1185** (analytics ping, PostHog) — Runway
removed analytics in #9.
- **openusage robinebers#1111 / robinebers#1136 / robinebers#1106** (scroll / Settings lag / SVG
parse) — Runway already has `ReorderFrameStore`, parsed-once
`ProviderMark`, and the rebuilt popover path. Taking their patch would
duplicate that work.
- **openusage robinebers#1137 / robinebers#1165 / robinebers#1141** (Codex Session default, Fable
order) — Runway already hides Codex Session by default and already
places Fable directly below Weekly. Layout defaults stay an owner
decision.
- **openusage robinebers#1134 Grok Bot meter** — new Cursor metric. AGENTS.md
requires owner confirmation of the four defaults before adding it; this
PR only takes the dashboard label rename and the `grok-bot-*` pricing
aliases.
- **openusage robinebers#1139** (Antigravity local spend) — new scanner, protobuf
decoder, and new metrics. Too large for this wave and needs the same
default-placement call.
- **openusage robinebers#1195** (OpenCode Codex OAuth attribution) — new scanner
sharing Codex request pricing. Real feature, own follow-up; folding it
in here would bloat the PR.
- **openusage robinebers#1164** (Claude multi-account) — Runway already discovers
Claude homes and gives each account its own card.
- **openusage robinebers#1177** (Codex fallback pricing Settings) — extra Settings
surface; earlier port waves skipped extra reset/settings chrome for the
same reason.
- **openusage robinebers#1179** (dead pin ID remap) — OpenUsage layout keys and
old Antigravity IDs. Runway installs never held those keys (different
defaults domain), and schema v3/v4 are already used for the beta-channel
and telemetry retirements.
- **openusage robinebers#1172** (bound log memory) — Runway already rejects
non-finite / overflowing token counts at the parse boundary instead of
clamping them.
- **openusage robinebers#1167 / robinebers#1016** (sub-1% "Not started", untouched pacing) —
already in Runway (`used <= 0`, `Pace.evaluate` returns nil when
unused).
- **openusage robinebers#1128** (Sparkle 2.9.6) — still a relevant bump; leaving
it to Dependabot rather than mixing a package-resolution change into
this accuracy PR.
- **openusage robinebers#1170 / robinebers#1159 / robinebers#1143 / robinebers#1163** (contribution policy,
screenshot assets, test-suite cleanup) — not user-facing on Runway, or
would churn tests without changing behavior.
- **openusage robinebers#1196** (legacy Codex iCloud identity) — Runway's sync
identity path is already fork-specific.

Gemini 3.8 Flash output is **$3.50**, from [Cursor's
table](https://cursor.com/docs/models-and-pricing.md), not upstream's
$3.75 Google API rate. That matches how Runway priced Gemini 3.7.

## Tests

- Desktop: prefixed key for the signed-in account wins; foreign `acct:`
keys are ignored; a scoped V2 tombstone suppresses the V1 alias;
`load()` reads `lastKnownAccountUuid`.
- Claude: no credentials plus local logs → spend tiles and Not logged
in, not an error card. Empty machine still errors.
- Grok: subagent ledger is included; fork replay of a shared prompt
still counts once.
- Codex: `self_serve_business_prolite` → Business Premium, weekly-only
window.
- Pricing: Astra / 3.8 Flash / Fable 5.1 / GLM 5.3 / grok-bot slugs and
router labels resolve. `testEveryAliasCanonicalResolves` covers the new
rules.
- Cursor mapper tests updated to the new labels; widget IDs unchanged.

`swift test --filter
"ClaudeDesktopAuthStoreTests|ClaudeProviderTests|CursorProviderTests|CursorUsageSummaryTests|GrokLogUsageScannerTests|CodexUsageMapperTests|PricingBundledResourceTests|LayoutStoreTests"`
— 179 tests, 1 skipped, 0 failures.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant